Meta's First Real-Time Audio Perception Model Arrives: Transcribe 20 People Speaking Simultaneously, $3 per 1000 Minutes
Meta has launched its first real-time audio perception model, Muse Voice Transcribe, which can be called via Model API, priced at $3 per 1000 minutes (approximately $0.18 per hour). The model integrates streaming speech recognition, speaker separation, and endpoint detection, supporting real-time transcription without waiting for the entire audio, and automatically distinguishing more than 20 speakers for track-by-track output.